前兩天介紹了 GraphRAG,今天來看另一個比較輕量的實作:LightRAG。
Microsoft GraphRAG 的想法很完整,但 indexing 階段要做 Entity / Relationship Extraction、Community Detection、Community Report,Global Search 還會搭配 Map-Reduce,整體流程和成本都不算小。
LightRAG 的方向比較直接:保留 Knowledge Graph,但把 Retrieval 設計得更輕量,並同時保留傳統 Text Chunk Vector Search。官方目前把它定位成一套 lightweight knowledge-graph RAG framework,並使用 Graph + Vector 的雙層架構。
LightRAG 的 indexing 邏輯跟前面介紹的 GraphRAG 一樣,會先對文件做 Chunking,再利用 LLM 從 Chunk 中抽取:
Entity
Relationship
例如:
OpenAI develops AI models.
Microsoft invested in OpenAI.
可能得到:
OpenAI
Microsoft
AI Model
Microsoft ── invested_in ── OpenAI
OpenAI ── develops ── AI Model
目前 LightRAG 對 Entity 會保存名稱、類型與 description;Relationship 則會保存 source、target、relationship keywords 與 description。
這些資料最後不只會形成 Knowledge Graph,也會建立 Vector Representation。
LightRAG 實際上同時需要幾種不同 Storage:
KV Storage
Vector Storage
Graph Storage
Document Status Storage
其中 Vector Storage 會保存 Text Chunk、Entity、Relationship 的向量,Graph Storage 則保存 Entity 與 Relationship 形成的 Knowledge Graph。官方預設提供簡單的本地 storage,也可以換成 PostgreSQL、Milvus、Qdrant、Neo4j 等後端。
所以它的 indexing 可以先簡化成:
Document
↓
Chunk
↓
LLM Extraction
↓
Entity + Relationship
↓
Knowledge Graph
同時:
Chunk / Entity / Relationship
↓
Embedding
↓
Vector Storage
LightRAG 論文最主要的概念之一是 Dual-Level Retrieval。
Query 進來之後,LightRAG 會先從 Query 中抽出兩種 Keyword:
Low-level Keywords
High-level Keywords
Low-level 比較偏具體 Entity、名稱或細節,例如:
OpenAI
Sam Altman
Microsoft
GPT-5
High-level 則比較偏概念與主題,例如:
AI investment
corporate partnership
large language models
目前官方實作仍然會先利用 LLM 從 Query 中抽 low_level_keywords 與 high_level_keywords。Low-level keywords 主要用來找 Entity;High-level keywords 主要用來找 Relationship。
所以可以理解成:
Query
↓
Keyword Extraction
↓
├─ Low-level Keywords
│ ↓
│ Entity Vector Search
│
└─ High-level Keywords
↓
Relationship Vector Search
這跟 Microsoft GraphRAG Local Search 一開始只從 Entity 找 semantic seed 的做法有點不同。
LightRAG 不只搜尋 Entity,也直接替 Relationship 建 Vector Index,因此 Query 可以從:
具體 Entity
和:
抽象 Relationship / Theme
兩個方向同時進入 Knowledge Graph。
LightRAG 裡 local 和 global 的意思,跟 Microsoft GraphRAG 不完全一樣。
local 主要偏 Entity:
Low-level Keywords
↓
Entity Vector Search
↓
相關 Entity
↓
取得 Entity 周圍的 Relationship
↓
相關 Text Chunks
適合:
某個人是誰?
某家公司做了什麼?
A 跟 B 有什麼直接關係?
官方目前也把 top_k 在 local mode 定義成主要取得 Entity 的數量。
global 則偏 Relationship:
High-level Keywords
↓
Relationship Vector Search
↓
相關 Relationship
↓
取得相關 Entity / Text Chunks
適合比較宏觀、跨 Entity 的問題,例如:
這些公司之間主要有哪些合作模式?
AI 產業中有哪些重要的投資關係?
這點跟 Microsoft GraphRAG 的 Global Search 差異很大。
Microsoft GraphRAG 的 Global Search:
Community Reports
→ Map-Reduce
LightRAG 的 Global Retrieval 則主要還是:
Relationship Vector Search
→ Graph Context
所以 LightRAG 沒有必要先做 Community Detection 和 Community Report,這也是它相對輕量的一個原因。
hybrid 很直觀,就是把 Local 和 Global 都做:
Low-level Keywords
→ Entity Retrieval ─────┐
├→ Merge Context
High-level Keywords │
→ Relationship Retrieval┘
也就是同時看:
具體 Entity
+
整體 Relationship
這也是原始 LightRAG 論文裡 Dual-Level Retrieval 的主要概念。
LightRAG 現在還有兩個很值得注意的 Mode。
naive 就是傳統 RAG:
Query
↓
Chunk Vector Search
↓
Top-K Chunks
↓
LLM
完全不使用 Knowledge Graph。
而 mix 則會把:
Local Graph Retrieval
+
Global Graph Retrieval
+
Chunk Vector Retrieval
一起使用。
目前官方預設 Query Mode 已經是 mix,官方 README 也直接建議一般情況優先使用 mix,因為它同時保留 Knowledge Graph 與原始 Text Chunk Retrieval。
所以整體可以整理成:
Entity Retrieval ─────┐
/ │
Query → Keywords ├→ Context → LLM
\ │
Relation Retrieval ───┤
│
Query → Chunk Vector Search ──────────┘
這其實也剛好回到前幾天一直提到的問題:
Graph Retrieval 不一定要取代 Text Retrieval。
LightRAG 現在的 mix mode 就是直接把兩種 Retrieval 留著。
兩者最大的差異不是「有沒有 Knowledge Graph」,而是對 Graph 的使用方式不同。
Microsoft GraphRAG 在 indexing 階段會建立:
Entity
Relationship
Community
Community Report
Community 讓它可以做比較強的 Global Search,但相對需要更多 indexing 與 LLM processing。
LightRAG 則沒有走 Community Report 這條路,而是直接對:
Entity
Relationship
Text Chunk
都建立可以搜尋的 representation,再透過不同 Query Mode 組合 Retrieval。
可以粗略理解成:
Microsoft GraphRAG
→ 先把 Graph 整理成更高階的 Community Knowledge
→ Retrieval 架構比較重
LightRAG
→ 直接利用 Entity + Relationship + Chunk
→ Retrieval 架構比較輕
這也是 LightRAG 名字裡 Light 的主要精神之一。原始論文也把重點放在降低 graph indexing / retrieval 的複雜度,以及支援 incremental update。
LightRAG 另一個值得提的點是 Incremental Update。
Microsoft GraphRAG 這種需要 Community Detection、Community Report 的架構,新增資料後可能會影響 Community 結構。
LightRAG 的 Graph 相對單純,新文件進來後主要就是:
新 Chunk
↓
抽 Entity / Relationship
↓
Merge 到既有 Graph
↓
更新 Vector Storage
原始 LightRAG 論文也把 incremental update 當成主要設計目標之一,讓新的資料可以加入現有 Knowledge Graph,而不需要每次重新建完整 indexing pipeline。
這對實際 Knowledge Base 滿重要,因為 Production RAG 的文件通常不是一次匯入之後就永遠不變。
這篇如果時間不夠,我建議程式只放最小範例,不要深入 Storage。
from lightrag import LightRAG, QueryParam
初始化好 LLM、Embedding 與 Storage 後,把文件寫入:
await rag.ainsert(
"Microsoft invested in OpenAI. "
"OpenAI develops large language models."
)
Query 時切換 Mode:
result = await rag.aquery(
"What is the relationship between Microsoft and OpenAI?",
param=QueryParam(mode="local"),
)
也可以:
QueryParam(mode="global")
QueryParam(mode="hybrid")
QueryParam(mode="naive")
QueryParam(mode="mix")
目前 QueryParam 正式支援 local、global、hybrid、naive、mix 與 bypass,預設是 mix。
文章如果只是介紹架構,我會到這裡就停,不要再展開完整 LLM / embedding provider 設定,不然主題很容易變成 LightRAG 安裝教學。
LightRAG 的優點很明顯:相比 Microsoft GraphRAG,它少了 Community Detection、Community Report 與 Global Search 的 Map-Reduce,Indexing 和 Query Pipeline 都比較容易理解,而且現在的 mix mode 也保留了傳統 Chunk Retrieval,不會完全依賴 Knowledge Graph。
不過它的核心問題仍然跟所有 GraphRAG 一樣:
Entity 有沒有抽對?
Relationship 有沒有抽對?
同一個 Entity 有沒有正確 Merge?
Graph 品質不好,後面的 Retrieval 再漂亮也沒用。
而且 LightRAG 官方目前也特別提醒,Entity / Relationship Extraction 本身對 LLM 能力有一定要求,Production 環境不能只把它當作「傳統 RAG 多一張 Graph」這麼簡單。
所以我會把 LightRAG 理解成:
不是把 RAG 改成 Graph Search,而是在傳統 Chunk Retrieval 之外,再建立一條 Entity / Relationship Retrieval Path。
跟 Microsoft GraphRAG 相比,它選擇少做一些昂貴的高階 Graph Summarization,換取更簡單的 indexing、retrieval 與 incremental update。
再堅持一下,下班後就可以哭了。